Add B300 Dynamo+SGLang AgentX configs for DeepSeek V4 with aggregate DEP8 c384 / 添加含聚合式 DEP8 c384 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 - #3631
Conversation
新增基于 Mooncake 的 B300 DeepSeek V4 AgentX 配方,并适配 Python 启动器。
移除配方中与集群级每 GPU CPU 配置冲突的任务级 CPU 参数。
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…gentX The pluggable launcher resolves DeepSeek-V4-Pro-0813 on b300-dsxe to the shared /data/models copy for every non-vLLM framework, including multi-node Dynamo+SGLang, which previously read the node-local /scratch/models copy. With several points loading concurrently from shared storage, the TP8 c1 agg worker spent ~2.9h in weight loading and lost its etcd lease; the TP4 c4 agg worker never became healthy within the 4h window on the previous head. Pin every point of both configs to /scratch/models/DeepSeek-V4-Pro-0813 via the MODEL_PATH additional-setting, restoring the storage path under which this recipe set last passed. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…oint A point-level MODEL_PATH is opaque to the launcher, so srtctl's model preflight (srt-slurm v2.39.1) ran on the runner host and rejected the node-local /scratch/models path. Instead, add dynamo-sglang to the b300-dsxe override that already routes vLLM to DeepSeek-V4-Pro-0813@scratch: the checkpoint is then known to be node-local, model_paths maps the recipe alias to /scratch/models/DeepSeek-V4-Pro-0813, and preflight is skipped as it is for vLLM. Single-node SGLang keeps the shared /data/models copy. Drop the MODEL_PATH additional-settings added in the previous commit. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Replace the disaggregated 2P1D DEP8 c480 point with a one-node aggregate DEP8 c384 Dynamo+SGLang recipe that follows the single-node SGLang DEP8 c384 point: attention DP8 + EP8 Mega-MoE, FP4 indexer, prefill delayer (interval 20), chunked prefill 65536, max-running-requests 768, cuda-graph-max-bs-decode 544, mem-fraction-static 0.88, swa-full-tokens-ratio 0.075 and a HiCache ratio 3 write_back DRAM tier with the page_first_direct layout. The recipe uses this PR's lmsysorg/sglang:nightly-dev-20260916-c9a8fba9 image and its arg names (tp-size/dp-size/ep-size, attention-backend dsv4 instead of the deprecated compressed alias, no deprecated SGLANG_ENABLE_UNIFIED_RADIX_TREE). The Dynamo frontend tokenizes and routes, so SGLang-server-only settings (tokenizer workers, parsers, chat template, server warmup, keep-alive) are dropped; per-DP-rank KV events feed the Dynamo KV router. Drop enable-w4a4-mxfp4-megamoe from every recipe, so Mega-MoE runs its default FP8xFP4 kernels. 将 B300 分离式 2P1D DEP8 c480 测试点替换为单节点聚合式 DEP8 c384 Dynamo+SGLang 配方,沿用单节点 SGLang DEP8 c384 测试点:注意力 DP8 + EP8 Mega-MoE、FP4 indexer、prefill delayer(间隔 20)、chunked prefill 65536、max-running-requests 768、cuda-graph-max-bs-decode 544、 mem-fraction-static 0.88、swa-full-tokens-ratio 0.075,以及 HiCache 比例 3、write_back、page_first_direct 布局的 DRAM 层。 配方使用本 PR 的 lmsysorg/sglang:nightly-dev-20260916-c9a8fba9 镜像及其参数 命名(tp-size/dp-size/ep-size、以 attention-backend dsv4 替代已弃用的 compressed 别名、不再设置已弃用的 SGLANG_ENABLE_UNIFIED_RADIX_TREE)。由于 Dynamo 前端负责分词与路由,移除仅适用于 SGLang 服务端的设置(tokenizer worker、解析器、聊天模板、服务端预热、keep-alive);各 DP rank 的 KV 事件 供 Dynamo KV 路由器使用。 所有配方移除 enable-w4a4-mxfp4-megamoe,Mega-MoE 使用默认的 FP8xFP4 kernel。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
将 B300 Dynamo+SGLang AgentX 变更日志条目的 pr-link 指向 #3631。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
| cpus-per-gpu: 24 | ||
| salloc-args: [--mem=0] | ||
| volumes: | ||
| aiperf-cache: {path: /data/home/sa-gha-runner/aiperf-cache} |
There was a problem hiding this comment.
🟣 The new aiperf-cache volume for b300-dsxe is never mounted into the multi-node jobs this PR adds, so the PR's claimed "persistent AgentX and Hugging Face caches" don't actually persist. SRT_LANES[("b300-dsxe", LaunchPath.SRT_MULTI)] in infx/launch/drivers/srt/lanes.py:70 is SrtLane(frameworks=_DYNAMO) with no mounts=_AGENTIC_CACHES, so lane_mounts() (drivers/srt/config.py:187-197) never binds aiperf-cache or hf-hub-cache for this cluster even though the new agentic-coding recipes set AIPERF_DATASET_MMAP_CACHE_DIR=/aiperf_mmap_cache and HF_HUB_CACHE=/hf_hub_cache. Fix: add mounts=_AGENTIC_CACHES (or equivalent) to the b300-dsxe SRT_MULTI lane so the new volume is actually wired to agentic multi-node jobs. Pre-existing gap in lanes.py, but this PR's 5 new recipes plus the runners.yaml volume addition substantially widen who hits it.
Why this was flagged
Trigger: any of the 5 new dsv4 AgentX recipes launches on b300-dsxe via the SRT_MULTI path with IS_AGENTIC=1 (set because scenario-type is agentic-coding, per infx/matrix/plan.py:131). The job's container env sets AIPERF_DATASET_MMAP_CACHE_DIR=/aiperf_mmap_cache and HF_HUB_CACHE=/hf_hub_cache (e.g. agg-tp8-c1-mtp.yaml:127-128), but lane_mounts() only mounts volumes listed in SrtLane.mounts, and lanes.py:70 lists none. On the base branch the same gap exists for the one prior qwen3.5 disagg point; this PR adds the aiperf-cache volume definition to runners.yaml (implying it should now work) and 5 more call sites, without updating lanes.py, so the paths stay ephemeral container-local directories instead of host-backed caches, re-downloading/reprocessing data every run with no persistent cache or error to flag it.
Verification: nit. The aiperf-cache volume added at runners.yaml:581 (and pre-existing hf-hub-cache at :583) is never mounted into b300-dsxe multi-node jobs. The b300-dsxe srt-slurm block has no volume-mounts key, and lanes.py:70 is ("b300-dsxe", LaunchPath.SRT_MULTI): SrtLane(frameworks=_DYNAMO) with no mounts=_AGENTIC_CACHES. lane_mounts() (config.py:187-194) iterates only lane.mounts.
Set cuda-graph-max-bs-decode to 96, the per-DP-rank request cap (max-running-requests 768 / dp-size 8). SGLang already clamps captured decode batch sizes to the per-rank request pool, so 544 never captured more than this; the explicit value states the intended limit. 将 B300 聚合式 DEP8 c384 配方的 cuda-graph-max-bs-decode 设为 96,即每个 DP rank 的请求上限(max-running-requests 768 / dp-size 8)。SGLang 本就会把 捕获的 decode batch size 限制在每个 rank 的请求池以内,544 实际上从未捕获 超过该值;显式设置可明确预期的上限。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
The c384 warmup burst hit Dynamo's 5 s default request-plane ack timeout while the worker ingested the large AgentX requests; the frontend then marked the only worker unreachable and returned 500/503 for every warmup request. Set DYN_TCP_REQUEST_TIMEOUT to 60 s, as in the disaggregated recipes of this PR. c384 预热突发请求期间,worker 接收大量 AgentX 大请求时触发了 Dynamo 请求平面 默认 5 秒的确认超时;前端随后将唯一的 worker 标记为不可达,所有预热请求均返回 500/503。与本 PR 的分离式配方一致,将 DYN_TCP_REQUEST_TIMEOUT 设为 60 秒。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36974436489 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36974436489 |
|
/use 36845004131 |
|
@hshrivastava-droid staged run 36845004131: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-01~r36845004131 This run remains available across future |
Fill the interactivity gap between agg TP4 c4 (~244 tok/s/user) and disagg 1P1D DEP8 c64 (~128 tok/s/user) with two points: - agg TP4 c8: the TP4 c4 recipe at max-running-requests 16 and cuda-graph-max-bs-decode 16, plus a HiCache DRAM tier (ratio 2.75, write_through, direct I/O, page_first_direct, as in the B200 TP8 AgentX recipe on the same image) because TP4 c4 already fills about 60% of the GPU KV pool. - disagg 1P2D DEP8 c64: the 1P1D c64 recipe with a second decode worker, which halves each decode rank's load. 在聚合式 TP4 c4(约 244 tok/s/user)与分离式 1P1D DEP8 c64(约 128 tok/s/user)之间增加两个测试点: - 聚合式 TP4 c8:基于 TP4 c4 配方,max-running-requests 16、 cuda-graph-max-bs-decode 16;由于 TP4 c4 已占用约 60% 的 GPU KV 池, 增加 HiCache DRAM 层(比例 2.75、write_through、direct I/O、 page_first_direct,与同一镜像上的 B200 TP8 AgentX 配方相同)。 - 分离式 1P2D DEP8 c64:即 1P1D c64 配方增加第二个 decode worker,使每个 decode rank 的负载减半。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…o-sglang-agentx-dep8-agg Resolve the append-only perf-changelog.yaml conflict by keeping main's bytes unchanged and re-appending this PR's entry at the end. Co-Authored-By: Claude Opus 5.5 <[email protected]>
| python3 -m pip uninstall --break-system-packages -y mooncake-transfer-engine-cuda13 mooncake-transfer-engine-efa-cuda13 | ||
| python3 -m pip install --break-system-packages --no-deps mooncake-transfer-engine-efa-cuda13==0.3.13.post1 |
There was a problem hiding this comment.
can we get this in upstream?
There was a problem hiding this comment.
@functionstackx This is already covered by SGLang’s official AWS EFA instructions. They explicitly prescribe replacing the standard CUDA 13 Mooncake package with mooncake-transfer-engine-efa-cuda13==0.3.13.post1 using --no-deps.
This PR doesn't make any other changes to the container. I can submit a waiver explaining/documenting it, if it helps.
[by Claude Code]
Adds DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX recipes. This PR replaces #3190. It carries the same recipes, rebased on current
main, with these changes: a one-node aggregate DEP8 c384 point replaces the 2P1D DEP8 c480 point, two points (agg TP4 c8 and disagg 1P2D DEP8 c64) fill the interactivity gap between TP4 c4 and 1P1D c64, and no recipe enables the W4A4 MXFP4 Mega-MoE path.enable-w4a4-mxfp4-megamoeis removed from every recipe. Mega-MoE (moe-a2a-backend: megamoe) uses its default FP8xFP4 kernels.mooncake-transfer-engine-efa-cuda13==0.3.13.post1wheel during disaggregated container setup; no SGLang source patch is applied./opt/amazon/efaor/opt/amazon/ofi-ncclmounts.b300-dsxefrom the node-local NVMe copy (DeepSeek-V4-Pro-0813@scratch, as for vLLM).Aggregate DEP8 c384
Ported from the single-node SGLang DEP8 c384 point (
dsv4-fp4-b300-sglang-agentic-hicache-mtp,benchmarks/single_node/srt-slurm-recipes/dsv4/sglang/b300-fp4-mtp/agentic.yaml,override_dep8_c384):prefill-decode-interval: 20,stream-interval: 20,chunked-prefill-size: 65536withSGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK=8320,max-running-requests: 768(96 per DP rank, twice concurrency),mem-fraction-static: 0.88,swa-full-tokens-ratio: 0.075, HiCache ratio 3 (write_back,direct,page_first_direct), DSpark K=6.lmsysorg/sglang:nightly-dev-20260916-c9a8fba9,tp-size/dp-size/ep-size,attention-backend: dsv4(compressedis a deprecated alias in this image), and noSGLANG_ENABLE_UNIFIED_RADIX_TREE(deprecated; the unified radix tree is now the default). Thinking is enabled throughSGLANG_DEFAULT_THINKINGandSGLANG_DSV4_REASONING_EFFORT, as in the other recipes.cuda-graph-max-bs-decode: 96, the per-rank request cap (768 / 8). The single-node recipe used 544, but SGLang clamps captured decode batches to the per-rank request pool, so 544 never captured more than this.X-Dynamo-Session-ID, and each DP rank publishes KV events (kv-events-config), so the router sees per-rank prefixes. One frontend; etcd and NATS run on the worker node. The frontend setsDYN_TCP_REQUEST_TIMEOUT: "60", as the disaggregated recipes do, so the c384 warmup burst does not exceed Dynamo's 5 s request-plane ack timeout.tokenizer-worker-num, tool-call and reasoning parsers, chat template,skip-server-warmup,SGLANG_TIMEOUT_KEEP_ALIVE.dram-utilization: 0.80matches the disaggregated config.Gap-filling points: agg TP4 c8 and disagg 1P2D DEP8 c64
Two points fill the interactivity gap between agg TP4 c4 (~244 tok/s/user) and 1P1D DEP8 c64 (~128 tok/s/user):
max-running-requests: 16andcuda-graph-max-bs-decode: 16. TP4 c4 already fills about 60% of the GPU KV pool, so this point adds a HiCache DRAM tier with the same settings as the B200 TP8 AgentX recipe on this image (hicache-ratio: 2.75,write_through,direct,page_first_direct).Local validation
validate_perf_changelogagainstmain; the changelog change is append-only.validate_config_fileandsrtctl migrate(already current) for all seven recipes.infx/tests/launch,infx/tests/clusters,infx/tests/matrix,infx/tests/srt_slurm, and the changelog workflow tests (769 passed); Ruff onmodels.py.c9a8fba9).AI model disclosure
claude-opus-5-5, via Claude Code): rebased the Add Mooncake B300 Dynamo+SGLang AgentX configs for DeepSeek V4 / 添加基于 Mooncake 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 #3190 commits onto currentmain, replaced the 2P1D point with the aggregate DEP8 c384 recipe, removed the W4A4 MXFP4 Mega-MoE flag, ran the validation, and prepared this PR.中文
新增 DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX 配方。本 PR 取代 #3190:包含相同的配方并 rebase 到最新
main,另有以下改动:以单节点聚合式 DEP8 c384 测试点取代 2P1D DEP8 c480 测试点;新增聚合式 TP4 c8 与分离式 1P2D DEP8 c64 两个测试点,填补 TP4 c4 与 1P1D c64 之间的交互性空档;所有配方均不启用 W4A4 MXFP4 Mega-MoE 路径。enable-w4a4-mxfp4-megamoe,Mega-MoE(moe-a2a-backend: megamoe)使用默认的 FP8xFP4 kernel。mooncake-transfer-engine-efa-cuda13==0.3.13.post1wheel,不修改 SGLang 源码。/opt/amazon/efa或/opt/amazon/ofi-nccl。b300-dsxe上 Dynamo+SGLang 的 DeepSeek-V4-Pro-0813 权重从节点本地 NVMe 副本(DeepSeek-V4-Pro-0813@scratch,与 vLLM 相同)加载。聚合式 DEP8 c384
移植自单节点 SGLang DEP8 c384 测试点(
dsv4-fp4-b300-sglang-agentic-hicache-mtp,benchmarks/single_node/srt-slurm-recipes/dsv4/sglang/b300-fp4-mtp/agentic.yaml,override_dep8_c384):prefill-decode-interval: 20的 prefill delayer、stream-interval: 20、chunked-prefill-size: 65536配合SGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK=8320、max-running-requests: 768(每个 DP rank 96,为并发数两倍)、mem-fraction-static: 0.88、swa-full-tokens-ratio: 0.075、HiCache 比例 3(write_back、direct、page_first_direct)、DSpark K=6。lmsysorg/sglang:nightly-dev-20260916-c9a8fba9、tp-size/dp-size/ep-size、attention-backend: dsv4(该镜像中compressed为已弃用别名),并且不再设置SGLANG_ENABLE_UNIFIED_RADIX_TREE(已弃用,unified radix tree 现为默认)。与其他配方一样,通过SGLANG_DEFAULT_THINKING和SGLANG_DSV4_REASONING_EFFORT启用思考模式。cuda-graph-max-bs-decode: 96,即每个 rank 的请求上限(768 / 8)。单节点配方使用 544,但 SGLang 会把捕获的 decode batch 限制在每个 rank 的请求池以内,因此 544 实际上从未捕获超过该值。X-Dynamo-Session-ID保持亲和,各 DP rank 发布 KV 事件(kv-events-config),使路由器能看到每个 rank 的前缀。使用单个前端,etcd 与 NATS 运行在 worker 节点上。前端与分离式配方一样设置DYN_TCP_REQUEST_TIMEOUT: "60",避免 c384 预热突发请求超过 Dynamo 请求平面默认 5 秒的确认超时。tokenizer-worker-num、工具调用与推理解析器、聊天模板、skip-server-warmup、SGLANG_TIMEOUT_KEEP_ALIVE。dram-utilization: 0.80与分离式配置一致。补充测试点:聚合式 TP4 c8 与分离式 1P2D DEP8 c64
以下两个测试点填补聚合式 TP4 c4(约 244 tok/s/user)与 1P1D DEP8 c64(约 128 tok/s/user)之间的交互性空档:
max-running-requests: 16、cuda-graph-max-bs-decode: 16。由于 TP4 c4 已占用约 60% 的 GPU KV 池,本测试点增加 HiCache DRAM 层,设置与该镜像上的 B200 TP8 AgentX 配方相同(hicache-ratio: 2.75、write_through、direct、page_first_direct)。本地验证
main运行validate_perf_changelog;changelog 改动仅为追加。validate_config_file与srtctl migrate(已是最新 schema)。infx/tests/launch、infx/tests/clusters、infx/tests/matrix、infx/tests/srt_slurm及 changelog 工作流测试(769 项通过);对models.py运行 Ruff。c9a8fba9)核对 SGLang 参数与环境变量名称。AI 模型披露
claude-opus-5-5,通过 Claude Code):将 Add Mooncake B300 Dynamo+SGLang AgentX configs for DeepSeek V4 / 添加基于 Mooncake 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 #3190 的提交 rebase 到最新main,以聚合式 DEP8 c384 配方取代 2P1D 测试点,移除 W4A4 MXFP4 Mega-MoE 参数,完成验证并准备本 PR。🤖 Generated with Claude Code